01 The Big Picture
Context windows went from 2K tokens (GPT-3) to 1M+ in four years. That 500× jump did not come from bigger GPUs alone — it came from a stack of architectural tricks, each relaxing a different constraint.
Recall the layering of this series. Doc 03 established that context is the currency you spend; doc 09 attacked the problem from above — skills loaded on demand, memory banks, subagents; doc 11 added retrieval. All of those run on top of a fixed model substrate with a hard window. This doc is about that substrate: what physically limits s (the sequence length), which architectural moves raise each limit, and how to choose between long context, retrieval, and agentic memory when you design a system.
The punchline up front: the context window is one number, but it encodes three separate problems — and the techniques below solve them one at a time.
02 What: the Window Is Not Memory
A "1M-token context" claim hides three distinct problems that scale differently:
① Fit
Can the tokens physically exist? KV cache for s tokens ≈ 2·layers·heads·dhead·s·bytes — at 128K tokens this is many GB (doc 22). Fit is a memory capacity problem.
② Attend
Can the model actually use them? Attention must find the relevant 0.1% among a million tokens — in one forward pass, with no rehearsal. Attend is a compute + quality problem.
③ Persist
Does anything survive the request? The window is wiped per request; even a cached prefix expires in minutes (doc 07). Persist is a state across sessions problem — and architecture does not solve it at all.
03 Why: the Quadratic Wall
Everything in this doc traces back to one formula. Self-attention compares every token with every other token, so prefill cost in one layer is:
Quadratic means doubling the window quadruples the attention compute. Going 32K → 128K is a 16× jump in attention FLOPs per layer, and 4× the KV memory. The rest of the transformer is only O(s·d²) — so beyond roughly s ≈ d, attention itself is the dominant cost and the wall every "long context" technique attacks:
| s (tokens) | Relative attention FLOPs (vs 8K) | What breaks first |
|---|---|---|
| 8K | 1× | nothing — comfortable |
| 128K | 256× | KV cache exceeds GPU memory (doc 22) |
| 1M | ~15,600× | compute + bandwidth + positional generalization, simultaneously |
And there is a second, quieter asymmetry: capacity is cheap but fidelity is scarce. Fitting 1M tokens is a storage question; using them is an attention question. The next sections attack cost first (positional math, sparsity, parallelism), then measure the fidelity tax that remains (lost in the middle), then step outside the window entirely (RAG, agentic memory) because some fidelity can only be bought with retrieval.
Capacity — an engineering problem
Grows solvable: divide KV across devices (ring attention), sparsify the pattern (O(s·√s)), or stream with constant state (O(s)). Every adversary here has a counter-move.
Fidelity — a physics problem
One forward pass, no rehearsal, soft selection among s candidates. Positional math and sparsity buy back only part of it — the residual is the U-curve. Fidelity is ultimately rationed.
04 How — One Question Searching for Its Needle
One question ("what was the refund clause?") buried in a huge corpus. Step through the four architectures that try to connect them.
Steps 1–2 stay inside the window and pay quadratic + fidelity costs. Steps 3–6 move the problem outside: pay retrieval + distillation cost instead, and keep the window small. The rest of this doc prices each lane precisely.
05 Positional Math — Stretching RoPE, Bending ALiBi
Models learn position encodings only up to Ltrain. Tokens beyond that arrive with positions the rotary embedding (RoPE, doc 02) has never seen — attention quality collapses not because weights fail, but because rotation frequencies are out of distribution. The fixes interpolate instead of extrapolate:
06 Attention-Preserving Approximations — Pay for Less
If full attention is O(s²), the first family of tricks keeps softmax attention but sparsifies the pattern:
Local + global hybrid
Most layers use a sliding window of w tokens; a few layers (usually every k-th) attend to everything. Longformer and modern open-weights models use this shape. The window layers are O(s·w).
Longformer / BigBird patterns
Sliding window + global tokens + random links. Provably Turing-complete, and O(s·(w + g + r)) ≈ O(s·√s) when w, g, r scale as √s — a 128K sequence costs like a ~1.6M-token full-attention pass at the same per-pair price.
Linear attention / SSMs
Replace pairwise comparison with a fixed-size recurrent state: O(s) time, O(1) state. Constant memory per token, but the state must compress everything it has seen. Full treatment: doc 25 — Linear Attention & State-Space Models.
The hybrid pattern's key question: if most layers see only w neighbors, how far can the model see? Receptive field compounds through depth:
07 Ring Attention — When Context Became a Cluster Property
Even with sparsity, KV cache for a 1M-token prompt doesn't fit on one GPU. Context parallelism (ring attention) splits the sequence across P devices and makes the s² communication pipelined instead of all-at-once:
This is why 1M-token contexts are a cluster property, not a model property: the number in the model card is a claim about a fleet — how many GPUs sit in the ring — as much as about weights. And it is why long-context input tokens are priced steeply: your prompt's prefill is spread across a ring of devices whose interconnect, not FLOPs, sets the floor (doc 08's bandwidth story, repeated at datacenter scale).
08 Lost in the Middle — the Fidelity Tax
Fix fit, fix cost — and a quality curve remains. Liu et al. (2023) probed models with a needle placed at varying depths in long contexts. Accuracy vs needle position is a U:
Empirical: retrieval accuracy vs needle position in a long context
Interpretation through everything so far: attention has no "rehearsal loop" — a token seen once at position 300,000 gets exactly one pass of soft selection among millions of candidates. Primacy and recency edges are protected by position-bias and by RoPE/ALiBi locality; the middle is where quadratic sparsity and positional blur bite hardest. Three engineering consequences:
Put the needle at the edges when you control layout: critical instructions at the start, the current question and top retrieved chunks at the end. Rerank so best chunks land last. Test with needles placed at 25/50/75% depth, not just at the end.
Assume "it fits" means "it's used." A 1M-token dump with the answer in the middle can underperform a 2K-token RAG prompt — capacity was bought, fidelity wasn't. Don't benchmark long context with the needle at one fixed position.
09 RAG as a Memory Hierarchy
Everything above stayed inside the window. Retrieval (doc 11) steps outside and treats the window as a cache tier: L1 = the context window (small, expensive, high fidelity), L2 = the vector store (vast, cheap, must be addressed by search). The whole design question becomes: what belongs in L1 this turn?
Retrieval score math
The top-k tradeoff
k is a precision/recall dial over context tokens, not documents: raise k and recall of the needle rises (good) but the prompt fills with near-miss filler that dilutes attention (the U-curve's middle), costs prefill FLOPs, and pushes the true answer deeper into "lost" territory. The optimum is usually k ≈ 3–8 good chunks, not k = 50 — recall at the retrieval stage, precision at the window stage.
10 Agentic Memory — Persisting Across Sessions
RAG persists documents. Agentic memory (doc 09's memory banks, hardened) persists the agent's own conclusions: decisions, failed approaches, user preferences — as files the agent reads and writes. Architecture matters here too, in two numbers.
Hierarchical summarization is nearly free
Pay n tokens once to compress; afterwards, any question walks root → section → detail, reading O(log n) tokens instead of O(n). This geometric series is why "summarize, then index summaries" beats both "keep everything" (quadratic attention, lost middle) and "summarize flat" (one lossy bottleneck).
Consolidation policy — when to compact
A session's raw transcript is like dirty cache lines (doc 07): rewriting history mid-conversation invalidates the KV prefix from the first change point — the full-price recomputation rule from doc 07 applies to your own memory edits. So consolidation is batched:
Write memory additively during the session (append-only notes at the tail — prefix preserved, cache stays warm). Compact into stable summary files only at session boundaries or large batch points. Tag each memory with scope + expiry ("project-X preference", "stale after deploy").
Rewrite the middle of a live conversation per turn — you pay full prefill on everything after the edit. Don't store raw transcripts as memory; store distilled claims. Don't let memory grow unbounded — retrieval precision decays like an unindexed cache.
11 The Decision Framework — RAG vs Long Context vs Agentic Memory
Three tools, three cost structures. Estimate tokens per query for a corpus of size N with a query that needs m tokens of truth:
| Long context (dump it all) | RAG (retrieve top-k) | Agentic memory (distill + recall) | |
|---|---|---|---|
| Tokens read / query | ≈ N (everything, every time) | ≈ k·c chunks + query (k·c ≪ N) | ≈ S summaries + fetched detail on demand |
| Cost shape | prefill O(N) + O(N²) attention, every turn | embed query + rerank + prefill O(k·c) | amortized: Σ n/2^ℓ once, then O(log N) reads |
| Quality / fidelity | best case — model sees raw text; worst case — lost-in-the-middle | bounded by retrieval recall; misses are silent | bounded by what was distilled; drift if never re-consolidated |
| Failure mode | cost explosion + attention dilution | the needle was never fetched (no error signal!) | confidently stale memory |
| Choose when | needle-in-haystack QA over one corpus; cross-document reasoning where relevance is unknown | latency-critical fresh facts; corpora larger than any window; per-turn cost must be flat | multiple sessions; agent must remember decisions and preferences |
Multiply the "tokens read" row by your price per token and you have cost per query. For N = 500K and k·c = 4K, RAG reads ~0.8% of what a long-context dump reads — but if retrieval recall for your needle is below ~90%, you silently answer worse. Hybrid is the honest default: memory files (who the user is, standing decisions) + RAG (corpus facts) + long context (single hard documents), each handling the failure mode of the others.
12 The Memory-Tier Diagram
Context as a storage hierarchy — size grows downward, fidelity per token falls
Same shape as doc 08's GPU hierarchy, one level of abstraction up: L1 is fast and tiny (registers ↔ window), L2 is the working set you page in (RAM ↔ vector store), L3 is durable and slow to reach (disk ↔ memory files). Every performance lesson transfers: keep the hot path small, page on demand, batch your write-backs.
13 Mental Models
The context window is RAM — every address must be resident to be used. RAG is the MMU: the program (model) references "all my documents" by logical address, and the system pages in only the working set. A retrieval miss is a page fault; a reranker is the replacement policy. Lets you reason about: why bigger "RAM" (window) doesn't remove the need for an MMU (retrieval) — it just changes the working-set policy.
Window = L1 (small, fastest, most expensive per byte); vector store = L2; memory files = L3/disk. Ring attention is NUMA: capacity lives across nodes, and locality of the data path (the KV ring) decides cost. Lets you reason about: why you tune k like an L1 size — big enough for the working set, small enough that eviction noise (filler chunks) doesn't thrash attention.
14 Common Misconceptions
"1M-token context means the model can reason over 1M tokens." It means the tokens fit (problem ①). Attention quality at distance is governed by positional scaling, sparsity patterns, and the U-curve — a needle at 60% depth may be less reliably used than in a 10K window.
"Long context killed RAG." The cost math says otherwise: reading 500K tokens per query vs 4K is a >100× prefill difference, quadratic attention on top, and zero freshness guarantees. Long context changed RAG's design point (bigger chunks, whole-document retrieval, fewer round trips) — it didn't replace the hierarchy.
"Linear attention / SSMs are just better transformers." O(s) cost is real, but a fixed-size state must forget. For pure recall tasks over long ranges, softmax attention's exact pairwise lookup still wins; hybrids (a few full-attention layers among SSM layers) exist precisely because neither alone dominates.
"Memory = a longer conversation history." Persisting raw history is quadratic growth with no abstraction. Useful memory is distilled (the Σ n/2^ℓ hierarchy), scoped, and consolidated in batches — otherwise you re-import the window's problems into permanent storage.